Repurposing Benchmark Corpora for Reconstructing Provenance

نویسندگان

  • Sara Magliacane
  • Paul T. Groth
چکیده

Provenance is a critical aspect in evaluating scientific output, yet, it is still often overlooked or not comprehensively produced by practitioners. This incomplete and partial nature of provenance has been recognized in the literature, which has led to the development of new methods for reconstructing missing provenance. Unfortunately, there is currently no agreed upon evaluation framework for testing these methods. Moreover, there is a paucity of datasets that these methods can be applied to. To begin to address this gap, we present a survey of existing benchmark corpora from other computer science communities that could be applied to evaluate provenance reconstruction techniques. The survey identifies, for each corpus, a mapping between the data available and common provenance concepts. In addition to their applicability to provenance reconstruction, we also argue that these corpora could be reused for other tasks pertaining to provenance.

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

Towards Reconstructing the Provenance of Clinical Guidelines

Understanding the provenance of clinical guidelines is important for both practitioners and researchers as it allows for deeper understanding of the provided recommendations and could potentially provide a basis for updating guidelines. Often such provenance is incomplete or unavailable. We describe a prototype of a multi-signal pipeline for reconstructing provenance and show preliminary result...

متن کامل

Reconstructing Provenance

Provenance is an increasingly important aspect of data management that is often underestimated and neglected by practitioners. In our work, we target the problem of reconstructing provenance of files in a shared folder setting, assuming that only standard filesystem metadata are available. We propose a content-based approach that is able to reconstruct provenance automatically, leveraging sever...

متن کامل

Reconstructing Provenance Preliminary Results - Technical Report

Therefore, we developed a complementary approach, which considers the simpler problem of reconstructing provenance intended as dependencies between entities. The rationale is that once we are able to distinguish dependent entities, it becomes possible to refine the dependency relationships into sequences of operations. In the following, we describe a prototype implementation of this approach an...

متن کامل

Automatic Metadata Annotation through Reconstructing Provenance

Annotating datasets with metadata is an important part of organizing and curating data. However, it is a time consuming process and often not done in a rigorous fashion. In this paper, we propose a new approach to annotating datasets through the use of reconstructed provenance. A detailed survey of the related work in this area is given. Additionally, we provide an overview of our approach for ...

متن کامل

Reconstructing Human-Generated Provenance Through Similarity-Based Clustering

In this paper, we revisit our method for reconstructing the primary sources of documents, which make up an important part of their provenance. Our method is based on the assumption that if two documents are semantically similar, there is a high chance that they also share a common source. We previously evaluated this assumption on an excerpt from a news archive, achieving 68.2% precision and 73...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

عنوان ژورنال:

دوره   شماره 

صفحات  -

تاریخ انتشار 2013